Back

Theoretical and Applied Genetics

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Theoretical and Applied Genetics's content profile, based on 49 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Early-life stage phenomic prediction of field agronomic traits across breeding cycles in intermediate wheatgrass

Harris, Z. N.; Braley, J.; Cassetta, E.; Crain, J.; DeHaan, L.; Van Tassel, D.; Miller, A.; Rubin, M. J.

2026-08-31 plant biology 10.64898/2026.08.28.747871 medRxiv
Top 0.1%
12.1%
Show abstract

Perennial grains represent a promising frontier for sustainable agriculture, but breeding progress is constrained by the accessibility of genotyping and the difficulty of evaluating complex traits expressed for multiple years after establishment across heterogeneous environments. Phenomic selection may help address these challenges by using inexpensive, scalable, high-dimensional phenotypes collected early in development, although the robustness of such predictions across breeding cycles remains uncertain. Here, we compared genomic selection and phenomic selection across two breeding cycles of Thinopyrum intermedium (intermediate wheatgrass; IWG; Kernza(R)), comprising approximately 2,280 individuals from maternal half-sib families evaluated across multiple field sites and years. We constructed relationship matrices from genomic markers and early-life stage phenomic data, including seed and leaf color (HSV), CropReporter multispectral reflectance and indices, and cycle-specific hyperspectral reflectance sensors. Genomic models provided the strongest predictions on average across all field traits in both cycles. Among phenomic predictors, leaf HSV was consistently the most informative, whereas CropReporter and hyperspectral data showed lower and more trait-dependent performance and seed HSV provided little predictive value. Genomic, leaf HSV, and CropReporter models transferred across breeding cycles with little apparent loss of predictive ability relative to within-cycle validation, demonstrating that their predictive signals were not restricted to a single breeding cycle. Early-life stage leaf HSV emerged as a practical, accessible tool for germplasm thinning and early-stage prioritization in perennial breeding programs. Despite limited similarity among relationship matrices, multi-relationship-matrix models rarely improved prediction beyond the stronger constituent single-relationship-matrix model. Together, these results show that early-life stage phenomic data provide reproducible information about agronomic performance expressed years later, but that predictor complexity and data integration do not guarantee improved prediction.

2
Yield losses associated with peanut smut incidence in Argentina: a quantitative synthesis across field studies

Cazon, L. I.; Gonzalez, N. R.; Del Ponte, E. M.; Costa de Carvalho, A. C.; Asinari, F.; Camiletti, B. X.; Paredes, J. A.

2026-08-11 plant biology 10.64898/2026.08.11.744131 medRxiv
Top 0.2%
5.4%
Show abstract

Peanut smut, caused by Thecaphora frezzii, is an important constraint to peanut production in Argentina, but quantitative estimates of yield losses across environments remain limited. We quantified the relationship between disease incidence and kernel yield using 922 observations from 26 field studies conducted in Cordoba, Argentina, between 2021 and 2025. Study-specific incidence-yield relationships were analyzed using linear regression, random-effects meta-analysis, and linear mixed-effects models. Peanut smut incidence was consistently associated with yield reduction across studies. The estimated damage coefficient ranged from 24.2 to 28.7 kg ha-{superscript 1} per 1% increase in disease incidence, corresponding to a relative yield reduction of 0.74-0.87% of attainable yield. In contrast, attainable yield varied markedly among studies, ranging from 1,370 to 5,409 kg ha-{superscript 1}. Although an exploratory segmented analysis suggested a breakpoint near 12% incidence, subsequent moderator analyses, study- specific regressions, and normalized response curves provided no evidence of a biologically meaningful change in the damage coefficient across incidence or yield classes. These results indicate that differences among environments were primarily associated with attainable yield rather than with changes in the magnitude of disease-associated yield loss. The resulting damage function provides a quantitative basis for yield-loss assessment and disease management in peanut.

3
A haplotype-based breeding framework for the precise pyramiding of elite QTL alleles: a lettuce case study

Tu, Z.; Luo, G.; Xiao, L.; Wei, M.; Zhang, J.; Wang, X.

2026-08-20 bioinformatics 10.64898/2026.08.12.744550 medRxiv
Top 0.2%
5.4%
Show abstract

The efficient pyramiding of favorable alleles underlying complex traits remains a major challenge in crop breeding as most quantitative trait loci (QTLs) have not been resolved to causal genes, limiting their direct application in marker-assisted breeding. Although haplotypes provide more informative genetic units than individual markers, existing haplotype-based studies have largely focused on genetic interpretation and elite haplotype discovery, whereas computational frameworks for translating haplotypes into breeding decisions remain limited. Here, we developed HAPBDB, a haplotype-guided breeding framework that directly translates regional haplotypes into parental selection, cross design, and elite QTL pyramiding, and applied it to a lettuce genomic breeding panel. HAPBDB accurately reconstructed functional haplotypes at known loci and resolved elite haplotypes for five major QTLs controlling flowering time and yield. Integrating haplotype information across loci enabled systematic identification of accessions carrying complementary elite haplotypes and rational design of crosses that maximized favorable haplotype accumulation while minimizing segregating loci. Experimental validation using QTL-specific molecular markers demonstrated concordance between predicted and observed multi-locus genotypes across all designed F hybrids. Our results demonstrated that regional haplotypes can serve as practical breeding units even when the underlying causal genes remain unknown, thereby enabling the direct utilization of genetically mapped QTLs for precision breeding. By bridging the gap between genomic discovery and practical breeding, HAPBDB provides a practical framework for converting genomic information into breeding decisions and accelerating precision improvement of complex traits.

4
Harnessing Vitis germplasm diversity to dissect and predict adventitious rooting traits in grapevine

Sharma, S.; Lupo, Y.; Munoz, J.; Cochetel, N.; Nunez, V.; Gaspar, A.; Torres-Lomas, E.; Cantu, D.; Diaz-Garcia, L.

2026-08-27 genetics 10.64898/2026.08.24.746882 medRxiv
Top 0.2%
5.0%
Show abstract

Adventitious root formation (ARF) is a critical trait for the cost-effective propagation of grapevines in commercial nurseries. Poor rooting ability can limit the use and adoption of new rootstocks derived from underutilized Vitis species, constraining breeding efforts largely to the traditional trio: Vitis riparia, V. rupestris, and V. berlandieri. Despite its agronomic relevance, the genetic basis of ARF remains poorly characterized across the broader Vitis genus. In this study, we evaluated 308 accessions representing 18 Vitis species over three growing seasons, quantifying rooting performance at two developmental stages, callus-stage and post-transplant, alongside root biomass, cutting weight, and a derived transplant-response index. We observed extensive phenotypic variation both within and across species, and species rankings depended on the trait considered. V. riparia, V. rupestris and V. californica ranked among the top five species for all four rooting traits, whereas V. cinerea and V. candicans ranked among the lowest for root weight and post-transplant rooting. V. arizonica and V. acerifolia rooted well at the callus stage but were intermediate after transplanting, and V. berlandieri was among the weakest at the callus stage yet intermediate for post-transplant rooting. Repeatability was moderate to high for root weight (0.74) and callus-stage rooting (0.66), and lower for post-transplant rooting (0.47), reflecting both genetic control and season-to-season variation. Between-species differences accounted for 68% of the genetic variance in callus-stage rooting but only 10% in cutting weight. Rooting was associated with the climate of each accession's wild site of origin: after removing differences among species, accessions originating from sites with lower dry-season precipitation rooted better and produced more root biomass. Genome-wide association analysis using 3.4 million SNPs identified 54 significant SNPs resolving into 18 independent loci across four traits, with root weight contributing 12 of them. Candidate genes in linkage with these loci include a mitogen-activated protein kinase, a SCARECROW-LIKE GRAS transcription factor, PASTICCINO1, expansin A1, an AP2/ERF-RAV1 transcription factor, a tandem array of caffeoyl-CoA O-methyltransferases, and several sugar, peptide and nitrate transporters, implicating auxin-linked cell proliferation, cell wall and lignin remodeling, and solute transport. Genomic and phenomic prediction models yielded moderate accuracies across traits and seasons; up to r = 0.67 for post-transplant rooting within a season and r = 0.65 for previously unevaluated accessions. Moreover, the integration of spectral and genotypic data further improved predictive performance. Prediction accuracy was essentially flat between 5,000 and 50,000 markers. This study establishes a foundational framework for the genetic improvement of grapevine rootstocks, promoting broader use of resilient, high-performing, and clonally-propagable germplasm in viticulture.

5
Genetic mapping and genomic prediction for agronomic, grain compositional, and sensing-enabled traits in a cowpea MAGIC population along an environmental gradient

Berlingeri, J. M.; Lo, S.; Riggs, M.; Yun, H.; Kamangir, H.; Ranario, E.; Uyehara, I. K.; Mayanja, I.; Lao, A.; Dramadri, I. O.; Ongom, P. O.; Boukar, O.; Palkovic, A.; Bailey, B. N.; Earles, J. M.; Huynh, B.-L.; Diepenbrock, C. H.

2026-08-10 genetics 10.64898/2026.08.04.742818 medRxiv
Top 0.2%
5.0%
Show abstract

Cowpea (Vigna unguiculata [L.] Walp.) is a resilient grain legume and an important global source of dietary protein, yet the genetic and environmental basis of phenological and canopy development, as well as grain composition, remains incompletely characterized across production environments. In this study, we evaluated a cowpea multi-parent advanced generation intercross (MAGIC) population along an environmental gradient in California (with contrasting daylengths, temperatures, and soil types) using agronomic, grain compositional, and uncrewed aerial vehicle (UAV) and rover-enabled phenotyping. Near-infrared spectroscopy (NIRS) enabled assessment of grain compositional traits, while sensing-enabled time-series imaging captured canopy and reproductive dynamics. Quantitative trait locus (QTL) mapping identified 267 QTL, and genome-wide association studies (GWAS) detected 1,973 marker-trait associations. Integrating QTL mapping and GWAS results identified two major genomic hotspots affecting multiple traits. A chromosome 9 hotspot (5.8-6.0 Mb) was associated with flowering time and co-localized with sensing-enabled measures of flower and pod counts, plant height, and vegetation fraction, indicating broad effects on phenological and canopy development. A chromosome 8 hotspot (37.3-37.9 Mb) contained co-localized signals for seed weight, protein, starch, phytate, and moisture. A total of 22 prioritized candidate genes were identified within these and other loci with multi-environment QTL and GWAS support. Genomic predictive abilities were moderate to high for most traits and scenarios, with multi-trait MegaLMM outperforming RR-BLUP. Together, these results define major genomic regions controlling cowpea phenology, canopy development, and grain composition, and provide targets and strategies for breeding cowpea cultivars with favorable and environmentally resilient productivity and grain composition. Significance StatementTo dissect the genetic basis of cowpea productivity, adaptation, and grain composition, and how performance for these traits varies and can be predicted across environments, we combined multi-environment phenotyping, including sensing of canopy and reproductive traits, with quantitative genetic analyses in a multi-parental population. We identified genomic hotspots for seed size/composition and reproductive phenology and an across-environment predictive advantage for multi-trait vs. single-trait genomic prediction. Overall, these findings support the comprehensive improvement of cowpea.

6
Genetic mapping of a spontaneous short-grain mutation reveals a novel loss-of-function allele of SRS3 in rice

Montiel, M.; Angira, B.; Richards, J.; Famoso, A. N.

2026-08-09 genetics 10.64898/2026.08.03.742661 medRxiv
Top 0.3%
3.9%
Show abstract

Spontaneous mutations are a rare but important source of novel genetic variation, yet their detection and characterization within active breeding programs are seldom documented at gene-level resolution. Grain size and shape are key determinants of rice quality, yield, and market classification. Here, we report the discovery and genetic characterization of a spontaneous short-grain (SG) mutation arising in the long-grain wild-type (WT) advanced breeding line RU2002174 from the LSU AgCenter Rice Breeding Program. The SG phenotype was first observed in 2019 and segregated in subsequent generations as a single recessive gene across both indica and japonica genetic backgrounds. Genetic mapping localized the mutation to a 41.6 kb interval on chromosome 5. Whole-genome sequencing identified a single candidate causal variant: a G[->]T transversion in exon 4 of SRS3 (Os05g06280), introducing a premature stop codon and resulting in a truncated protein. This allele was absent from representative U.S. breeding germplasm and the IRRI 3K SNP database, demonstrating that it represents a novel spontaneous loss-of-function allele of a previously characterized grain-size gene. These findings document the real-time emergence of functional genetic variation in elite rice germplasm and highlight the importance of monitoring off-types during seed increase and purification in breeding programs. They also provide additional insight into the role of kinesin-mediated cell elongation in determining rice grain architecture.

7
Pyramiding panicle-level heat avoidance and grain-level heat tolerance improves rice grain appearance under high-temperature grain filling

Fukuda, H.; Sakamoto, T.; Yonemaru, J.-i.; Ogawa, D.

2026-08-21 plant biology 10.64898/2026.08.20.745907 medRxiv
Top 0.3%
3.9%
Show abstract

High temperature during grain filling increases rice grain chalkiness and deteriorates grain appearance under climate warming. Although several loci that reduce chalkiness have been identified, breeding strategies that integrate grain level heat tolerance with panicle level heat avoidance remain limited. Here we characterized SL2033, a chromosome segment substitution line carrying a long IR64 derived segment on chromosome 10, and evaluated the combination of the chromosome 10 segment with Appearance quality of brown rice 1 (Apq1), a quantitative trait locus associated with reduced heat induced chalkiness that acts at the grain level. Compared with its recurrent parent Koshihikari, SL2033 had longer flag leaves, altered vertical plant architecture, and lower panicle temperature. Total starch and protein contents were comparable between the two genotypes, whereas RNAseq analysis of the developing endosperm identified specific differences in heat, stress, and cell wall related transcripts. In a two year field trial, a pyramided line combining the SL2033 derived segment with Apq1 had the highest proportion of perfect grains and lowest frequencies of multiple chalky kernel types during the year with hotter grain filling conditions, with no detectable yield penalty. The pyramided line combined longer flag leaves, as in SL2033, with shorter panicle exsertion, as in an Apq1 near isogenic line, and had the lowest panicle temperature among the tested genotypes. Time series unmanned aerial vehicle imaging also detected genotype dependent differences in plant height during early grain filling, supporting distinct temporal patterns of plant development among the lines. These findings demonstrate that pyramiding genetic loci that confer panicle level and grain level heat tolerance is a promising strategy for improving rice grain appearance under high temperature field conditions, which are becoming increasingly prevalent.

8
Genome-wide dissection of tillering responsiveness to neighbour proximity in sorghum

Riaz, A.; Pearson, S.; Hunt, C.; Sukumaran, S.; Tao, Y.; Cooper, M.; Hammer, G.; Mace, E.; Jordan, D.

2026-08-14 plant biology 10.64898/2026.07.08.737219 medRxiv
Top 0.3%
3.1%
Show abstract

Tillering plasticity is a key adaptive trait in sorghum influencing resource use efficiency via a plants ability to adjust branching to neighbour density. Neighbour detection through red:far-red (R:FR) light sensing regulates this plasticity. While molecular pathways regulating tiller outgrowth are partly known, the genetic architecture underlying density-responsive tillering has not been resolved in any grass species. A sorghum diversity panel (n = 895) was evaluated over two growing seasons (2023 and 2024) with plant spacing ranging from 5 to 60 cm. A linear mixed model incorporating neighbour distance and tiller counts estimated genotype-specific response. GWAS was conducted on isolated plants (no neighbours within 60 cm) and on estimated responsiveness to neighbours. GWAS identified 52 baseline tillering QTLs and 50 for spacing responsiveness, with 10 overlapping, suggesting shared genetic control. Comparison with 41 R:FR pathway candidate genes revealed enrichment in responsiveness QTLs (5/50, 10%) versus baseline (0/52, 0%) (Fishers exact test, P = 0.025). Our model identified 40 unique density-responsive tillering QTL regions. Reducing genotype response to neighbour absence could be a selection target to develop water-efficient sorghum varieties where controlled architecture may be more valuable than natural plasticity.

9
Genetic Diversity and Population Structure of Maize Doubled Haploid Lines from Drought and Low Nitrogen Tolerant Populations

Ehemba, G. L.; Ifie, B. E.; DAS, B.; Abu, P.; Adjei, E. A.; Ayenan, M. A. T.; Garcia-Oliveira, A.; Ribeiro, P.; Manilal, W.; Tongoona, P.; Danquah, E. Y.

2026-08-13 genetics 10.64898/2026.08.06.743182 medRxiv
Top 0.3%
2.7%
Show abstract

Understanding the genetic diversity and population structure of breeding materials is essential for developing stress-resilient cultivars. In tropical maize, where drought and low soil nitrogen (low N) severely limit productivity, continuous development of tolerant varieties remains a priority. This study assessed the genetic diversity and population structure of 250 doubled haploid lines (DHLs) derived from five drought- and low N-tolerant tropical populations. Genotyping was performed using mid-density DArTseq markers, yielding 3,305 high-quality SNPs for analysis. Results revealed a moderate level of diversity among the DHLs, with an average genetic distance of 0.39, a polymorphism information content (PIC) of 0.33, and a minor allele frequency (MAF) of 0.29. These values reflect substantial allelic variation, important for identifying complementary parental combinations in hybrid development. Discriminant analysis of principal components (DAPC) grouped the DHLs into five distinct clusters, largely corresponding to their source populations, although some admixture was observed. This indicates that while the genetic backgrounds of the source populations were mostly retained, recombination introduced useful variation. Overall, the clear population structure and high diversity observed among these DHLs provide a strong genetic foundation for future maize improvement. These lines represent valuable resources for heterotic group formation, hybrid development, and recurrent selection schemes aimed at enhancing drought and low nitrogen tolerance in tropical maize.

10
Flywheel Genomics: Simultaneous trait discovery and genetic gain in plant breeding

Rice, B.; Ogoe, E.; Charles, J. R.; Melgar, E.; Marla, S.; Felderhoff, T.; Fritz, A.; Morris, G.; Pressoir, G.

2026-08-09 genetics 10.64898/2026.08.03.742257 medRxiv
Top 0.4%
2.4%
Show abstract

Genomic mapping has yielded extensive catalogs of quantitative trait loci underlying agronomic traits, yet translating these discoveries into breeding gains remains inefficient. Here, we introduce Flywheel Genomics, a framework that integrates trait discovery directly within rapid cycling breeding populations. Using empirical data from a smallholder-oriented sorghum breeding program, we demonstrate that recurrent intermating and selection maintain genetic diversity, effective population size, and recombination while reducing confounding from plant height and maturity. Within this population, we resolve loci underlying simple adaptive and complex environmentally responsive traits and generate large segregating populations for mapping and near-isogenic lines for locus validation. We further demonstrate applicability in a public wheat breeding program, where known agronomic loci were readily detected. Simulations show that rapid cycling better preserves the population genetic properties required for Flywheel Genomics than conventional pure line development. By integrating discovery with improvement, Flywheel Genomics reframes breeding programs as engines of both crop improvement and genetic insight.

11
Comparative assessment of genomic, phenomic, and metabolomic prediction models in biparental grapevine breeding populations

Borrelli, C.; Delannoy, L.; Chepca, H.; Calcaterra, M.; Chedid, E.; Arnold, G.; Dumas, V.; Baltenweck, R.; Maia-Grondard, A.; Hugueney, P.; Merdinoglu, D.; Duchene, E.; Avia, K.

2026-08-18 genomics 10.1101/2025.10.24.684307 medRxiv
Top 0.4%
2.3%
Show abstract

Accelerating grapevine breeding for disease resistance and climate adaptation remains constrained by long generation cycles. We benchmarked genomic (SNP), phenomic (NIRS), and metabolomic (untargeted LC-MS) prediction for 24 agronomic traits in a biparental population phenotyped over three years. Seven statistical frameworks and four tissue x timepoint combinations (wood; vineyard leaves at budbreak and flowering; greenhouse leaves at flowering) were evaluated, together with feature-wise BLUPs across samples. Cross-year and cross-population analyses with two additional populations assessed temporal robustness and transferability. Genomic prediction was most accurate (up to r = 0.83), metabolomic prediction was intermediate (up to r = 0.59), and phenomic prediction was lowest (up to r = 0.39) despite its lower acquisition cost. Metabolite features were more heritable than NIR wavelengths, for which most unexplained variation remained residual under the fitted model. Multi-omics integration produced limited overall gains. These results support genomic selection as the primary approach, with metabolomic or phenomic screening considered only for traits and sampling designs that show reproducible predictive signal.

12
Near-infrared phenomic and genomic prediction for seed protein in winter legume white lupin (Lupinus albus L.): A utility comparison

Castillo, M. P.; Oyebode, O. G.; Lenahan, A.; Orloski, A.; Wolfe, M.

2026-08-11 genomics 10.64898/2026.08.05.743001 medRxiv
Top 0.4%
2.1%
Show abstract

White lupin (Lupinus albus L.) is a cool-season grain legume with seed crude protein of 33-47%, competitive with soybean (Glycine max L.) meal. It also fixes nitrogen and mobilizes soil phosphorus. Because soybean is a summer crop, white lupin can occupy Southeastern winter fields as a complementary protein source. Breeding for seed protein is limited by the cost and throughput of reference phenotyping. To determine how each is best deployed, we compared the utility of near-infrared spectroscopy (NIRS)-based phenomic selection with genomic selection based on 246,847 SNPs from low-pass, whole genome sequencing in a panel of Auburn University breeding lines and USDA National Plant Germplasm System germplasm. A handheld NIR calibration against Dumas reference protein reached screening-grade accuracy (R2 = 0.81). Under common cross-validation, phenomic predictive ability was 0.93 and genomic was 0.12. The low genomic value was consistent with moderate heritability (H2 = 0.33) and strong genotype-by-year interaction. Beyond predictive ability, NIRS recovered superior accessions the strictest selection intensity, and 40 to 60 reference assays sufficed to calibrate the model. Handheld NIRS is a low-cost tool for protein calibration and early-generation screening, while genomic prediction remains suited to parental selection, together supporting a complementary strategy for legume breeding Plain Language SummarySoybean meal is the main protein source for livestock and fish farms in the United States. Because soybean is a summer crop, many Southeastern fields sit idle or grow low-value cover crops in winter. White lupin, a cool-season legume whose seeds are as protein-rich as soybean meal, makes a good complementary winter crop: it yields high-protein grain while serving as a cover crop that fixes nitrogen and frees up soil phosphorus for later crops. In our early-stage lupin breeding program, measuring seed protein by standard lab methods is slow and costly. We built a calibration that lets a handheld scanner estimate protein from light, and compared it with predicting protein from the plants DNA. The scanner gave accurate, low-cost protein screening from only about 40-60 lab tests, while DNA-based prediction remains suited to guiding parent selection. Used together, these tools offer breeders a practical path to develop high-protein white lupin. Core ideasO_LIHandheld NIRS provides screening-grade prediction of white lupin seed crude protein. C_LIO_LISpectra carried more usable protein signal than markers by measuring seed chemistry directly. C_LIO_LINIRS and genomic prediction serve different stages of a white lupin breeding program. C_LIO_LIAbout 40 to 60 reference assays sufficed to calibrate NIRS to near-full accuracy. C_LI

13
From Field Photosynthesis to Genetic Architecture: Insights from the First Dedicated Photosynthesis Hackathon

Matuszynska, A.; Sansa, O.; Adekoya, F. J.; Akinyemi, O. O.; Anokye, E.; Bashir, O. B.; Boyny, Z. Z. F.; Chukwuka, M. K.; Corvest, E.; Dada, A. O.; DellAcqua, M.; Ehemba, G. L.; Finkbeiner, A. J.; Hamabwe, S.; Hodehou, D. A. T.; Kacheyo, O.; Kamfwa, K.; Mhango, K. J.; Abdullahi, W. M.; Munduwe, G.; Ntukidem, S.; Obisesan, O. K.; Odesina, I. S.; Ogechi, N.-U.; Olaoye, O. D.; Olayinka, M. M.; Osei-Bonsu, I.; Rilwan, K. O.; Stival, L.; Tehar, Z.; Tende, R. M.; To, J.; Ugochukwu, U. K.; Unger, A.; van Aalst, M.; Vrbic, D.; Zhang, C.; Theeuwen, T. P. J. M.; Kramer, D. M.; Kromdijk, J.

2026-08-17 plant biology 10.64898/2026.07.24.740625 medRxiv
Top 0.4%
2.1%
Show abstract

Photosynthesis is among the most consequential yet genetically complex traits in crop plants, and translating its natural variation into actionable genomic targets remains a central challenge for breeding climate-resilient varieties. To start addressing this, researchers are generating increasingly large, multi-environment field photosynthesis datasets. Yet, these data have been structurally under-analysed since their inception. Here we report the outcomes of the first dedicated hackathon focused on computational mining of such field data held in Accra, Ghana, in March 2026. Bringing together data scientists, plant physiologists, geneticists, and breeders from Europe and Africa, these interdisciplinary teams used photosynthetic data collected with hand-held fluorometers to genome-wide marker data across four crop species: cowpea (Vigna unguiculata), barley (Hordeum vulgare), common bean (Phaseolus vulgaris), and potato (Solanum tuberosum). Despite using different species and methods, independent teams identified the same three key findings. First, mechanism-informed feature engineering and dynamic modelling recover genetic signals that are not detected or discarded in standard analysis pipelines, resulting in traits with improved heritability and meaningful associations with yield. Secondly, machine learning methods proved effective at uncovering genetic associations, with temporally resolved features substantially outperforming single time-point measurements. Third, raw chlorophyll fluorescence and absorbance traces consistently contained more information and predictive power than the extracted parameters currently used. A defining feature of this event was having experimentalists and data scientists working together, enabling AI approaches to be grounded in domain knowledge and biological mechanisms rather than relying on data alone.

14
Multi-environment GWAS analysis for photosynthetic light use efficiency in Arabidopsis thaliana

Nguyen, T.-P.; Erol, N. O.; Flood, P. J.; Moreira, C. N.; Theeuwen, T. P. J. M.; Harbinson, J.; Aarts, M. G. M.

2026-08-07 genetics 10.64898/2026.08.03.742429 medRxiv
Top 0.4%
1.9%
Show abstract

Photosynthesis is acknowledged as a potential target to increase crop yield. Improved photosynthesis may be achieved by conventional breeding, exploiting the available natural genetic variation for photosynthesis traits. This approach is challenging for crops due to limitations in high-throughput photosynthesis phenotyping, the highly polygenic nature of photosynthesis, and its strongly dynamic response to environmental changes. Recent advancements in phenomics make accurate and detailed photosynthesis phenotyping more feasible, with the model species Arabidopsis thaliana paving the way for applications in crops. In this study, we examined photosynthesis parameters over time in the global Arabidopsis HapMap diversity panel exposed to three conditions: optimal nutrient supply, low phosphorus supply and low nitrogen supply. Combined with two previous studies on photosynthesis in response to low temperature, and to a one-step change in irradiance from low light to high light, five high-quality datasets were systematically analysed using the same approach (with one million-maker set, uni- and multi-variate analyses). Our findings emphasize the genetic complexity of photosynthesis, detecting hundreds of significant quantitative trait loci, only a small number of which are robust, and of which most are condition specific. Robust loci, found in multiple conditions, exemplify those suited for conferring higher all-round photosynthesis, and targets for marker-assisted selection, contributing to environmental resilience, while the multitude of small-effect conditional loci suggest that genomic selection approaches may be more suited to improve crop photosynthesis.

15
A naturally occurring frameshift mutation in the UNUSUAL FLORAL ORGANS gene associated with the marimo floral phenotype in gerbera

Hattori, T.; Shimada, R.; Nagakura, M.; Ando, R.; Isobe, S.; Tajima, N.; Hirakawa, H.; Shirasawa, K.; Tominaga, A.

2026-08-14 genetics 10.64898/2026.08.09.743735 medRxiv
Top 0.5%
1.7%
Show abstract

BackgroundThe capitulum of Asteraceae is a highly specialized inflorescence whose formation requires the coordinated regulation of multiple developmental processes, including floral organ identity and floral meristem determinacy. The LEAFY (LFY)-UNUSUAL FLORAL ORGANS (UFO) regulatory module is known to play an important role in flower development; however, naturally occurring mutations affecting this pathway have not been genetically characterized in gerbera (Gerbera hybrida). ResultsIn this study, we characterized a novel gerbera mutant identified during a commercial crossing program and named it marimo based on its green, spherical capitulum. Morphological observations revealed the repeated formation of secondary and tertiary floret-like organs within primary floret-like organs. Scanning electron microscopy showed that the epidermal structure of the green organs in marimo was similar to that of wild-type involucral bracts. RNA sequencing identified numerous differentially expressed genes between marimo and the wild type, and network and Gene Ontology analyses highlighted gene groups associated with flower development, reproductive organ differentiation, and tissue structure formation. RNA-seq analysis showed increased expression of LFY and reduced expression of GGLO1, a PISTILLATA/GLOBOSA-like B-class MADS-box gene, in the marimo mutant. RT-qPCR analysis of a segregating population further confirmed reduced GGLO1 expression in marimo-type individuals. In addition, a single-nucleotide deletion was identified in the coding region of UFO. This deletion was predicted to cause a frameshift and a premature stop codon. In selfed progeny of No. 251, the UFO genotype was fully associated with capitulum phenotype, and only individuals homozygous for the mutant allele exhibited the marimo phenotype. ConclusionsThese results indicate that the naturally occurring frameshift mutation in UFO is the strongest candidate variant underlying the marimo phenotype. RNA-seq analysis showed increased LFY expression and markedly reduced GGLO1 expression in the marimo mutant. Reduced activity of the LFY-UFO regulatory module may therefore have altered the expression of GGLO1 and other floral organ development-related genes despite the continued expression of LFY. These changes may have affected both floral organ identity and floral meristem determinacy, resulting in the formation of green involucral bract-like organs and the repeated production of floret-like organs. The marimo mutant provides a useful genetic resource for investigating capitulum development in Asteraceae and may also serve as breeding material for introducing novel ornamental traits into gerbera.

16
CannSelect: A High-Quality Genotyping Platform for Cannabis sativa

Wilkerson, D. G.; Stack, G. M.; Carlson, C. H.; Quade, M. A.; Dowling, C. A.; Toth, J. A.; Murdock, M. J.; Jasinski, J.; Stansell, Z. J.; McKay, J. K.; Smart, L. B.

2026-08-21 genomics 10.64898/2026.08.18.745408 medRxiv
Top 0.5%
1.6%
Show abstract

The field of genomics has enabled extraordinary progress in horticultural crop research. However, there is still a need for cost-effective, high-resolution technologies flexible to the diversity found in emerging crops. To this end, we introduce CannSelect, a high-quality genotyping platform for Cannabis sativa. Designed for use in diversity analyses and trait mapping, probe targets were selected from four genotyped diversity panels and a curated gene list. This platform has been used to effectively map day-neutrality in a segregating population to the Autoflower1 locus with average capture efficiencies of 88.5%. With broad genome coverage, demonstrated target specificity, and reproducibility, CannSelect is expected to perform well across the diversity of C. sativa. We describe the methodology used to design CannSelect v1.0 and performance metrics for testing capture efficiency and target alignment in diverse genome assemblies. The CannSelect platform represents a robust and scalable, genome-wide genotyping tool for C. sativa researchers and breeders.

17
A century of soybean breeding increased photosynthetic capacity but not NPQ relaxation

Pereira de Oliveira, L.; Attri, K.; Doran, L.; Leonelli, L. B.; Long, S. P.; Ainsworth, E.

2026-09-01 plant biology 10.64898/2026.08.28.747836 medRxiv
Top 0.5%
1.5%
Show abstract

Accelerating photoprotective regulation to improve carbon assimilation is a promising strategy to increase crop productivity. Although rapid non-photochemical quenching (NPQ) relaxation has been validated as a target through metabolic engineering, it remains unclear whether conventional breeding has improved this trait. Here, we investigated whether more than a century of soybean breeding enhanced NPQ relaxation alongside light-saturated carbon assimilation and seed traits. We evaluated a historical panel of 24 soybean genotypes across vegetative and reproductive developmental stages by integrating NPQ relaxation, gas exchange parameters, xanthophyll-cycle pigment profiles, expression of key photoprotective genes (VDE, PsbS, and ZEP), seed number and seed weight. NPQ relaxation parameters were not consistently associated with genotype release year, seed number, or seed weight at either developmental stage. The only exception was the amplitude of the rapidly relaxing NPQ component (AqE), which was negatively correlated with all three variables during the reproductive stage. In contrast, genotype release year was positively associated with maximum net CO2 assimilation rate (Amax), maximum carboxylation rate of Rubisco (Vcmax), maximum electron transport rate (Jmax), seed number, and seed weight, while Amax and Vcmax were positively correlated with seed number and seed weight. These findings indicate that the greater photosynthetic capacity of modern genotypes was not accompanied by faster photoprotective response. Thus, photoprotective regulation has not kept pace with gains in photosynthetic capacity under field conditions. We conclude that rapid NPQ relaxation remains an important target for synchronizing photoprotection with the high photosynthetic capacity of modern soybean lines.

18
Impact of Reduced Chlorophyll Levels in Leaves on Soybean Yield, Seed Composition, Pod/Seed Photosynthesis, and Chlorophyll Levels in Pod and Seed Tissues

Jones, S. I.; Stutz, S. S.; Atalay, E.; Wang, Y.; Ort, D. R.; Cho, Y. B.

2026-08-19 plant biology 10.64898/2026.08.14.744892 medRxiv
Top 0.5%
1.4%
Show abstract

Soybean, a widely cultivated leguminous crop valued for its protein, amino acids, and oil, faces the challenge of maintaining protein levels, which have an inverse correlation with yield. Reducing leaf chlorophyll levels could increase seed protein levels without compromising yield; however, this is yet to be tested. Therefore, to understand the impacts of low chlorophyll mutations on soybean yield and seed composition, we screened and compared 25 low chlorophyll soybean mutants to their 11 dark green parents. PI548210 (Lincoln mutant) demonstrates a higher concentration of protein without affecting yield compared to its dark green parent PI548362 (Lincoln), suggesting it as a good candidate for further large-scale field trials. PI547555 (Y11/y11, Clark mutant) demonstrates a lower concentration of oil without impacting yield, alongside lower gross photosynthesis, but with chlorophyll levels in the pod and seed tissues that are comparable to its dark green parent PI548533 (Clark). These findings are consistent with the oil concentration of the soybean being influenced by pod and seed photosynthesis, which is correlated with pod height and row spacing. Chlorophyll levels in the leaf do not necessarily correlate with those in the pod and seed of low chlorophyll mutants, possibly due to substantially lower expression of chlorophyll synthesis genes in the pod and seed. SIGNIFICANCEO_LIPI548210 (Lincoln mutant), one of twenty-five low chlorophyll soybean mutants, demonstrates a higher concentration of soybean protein without affecting yield compared to its dark green parent (Figure 1 and Table 1). C_LIO_LIPI547555 (Y11/y11, Clark mutant), a low chlorophyll soybean mutant, demonstrates a reduced concentration of soybean oil without impacting yield, alongside lower gross photosynthesis in pod and seed tissues compared to its dark green parent (Figures 3 and Table 2). These findings suggest that the oil concentration of the soybean is influenced by pod and seed photosynthesis, which is in turn influenced by pod height and row spacing (Figure 2). C_LIO_LIChlorophyll levels in the leaf do not necessarily correlate with those in the pod and seed of low chlorophyll mutants, possibly due to substantially lower expression of chlorophyll synthesis genes in the pod and seed (Figure 5-6). C_LI O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=84 SRC="FIGDIR/small/744892v1_fig1.gif" ALT="Figure 1"> View larger version (55K): org.highwire.dtl.DTLVardef@4282dcorg.highwire.dtl.DTLVardef@9d565forg.highwire.dtl.DTLVardef@1918292org.highwire.dtl.DTLVardef@1359b1_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Two low chlorophyll mutants are as healthy as their dark green parents. Lincoln and its low chlorophyll mutant, left; Clark and its low chlorophyll mutant, known as Y11/y11, right. It can be seen by eye that the plants have low chlorophyll (light green/yellow leaves) but a similar growth habit to their dark green parents. See Supplemental Figures 1-4 for contrast, where low chlorophyll mutants are stunted in growth compared to their dark green parents. C_FIG O_TBL View this table: org.highwire.dtl.DTLVardef@657ec9org.highwire.dtl.DTLVardef@166e75borg.highwire.dtl.DTLVardef@df23c7org.highwire.dtl.DTLVardef@1a60124org.highwire.dtl.DTLVardef@194ed96_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 1.C_FLOATNO O_TABLECAPTIONComparison of seed yield, weight, seed composition between low chlorophyll mutants and their dark green parents. ANOVA is used with linear mixed model (random effect = block, fixed effect = variety). Least squares mean is used to compare. For yield and seed composition, N=4 blocks. For leaf chlorophyll (SPAD), N=40. Yield is average yield per plant (g). n.s. = not significant. C_TABLECAPTION C_TBL O_FIG O_LINKSMALLFIG WIDTH=179 HEIGHT=200 SRC="FIGDIR/small/744892v1_fig3.gif" ALT="Figure 3"> View larger version (26K): org.highwire.dtl.DTLVardef@7a368aorg.highwire.dtl.DTLVardef@192b8f0org.highwire.dtl.DTLVardef@1abb738org.highwire.dtl.DTLVardef@89e978_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 3.C_FLOATNO Light response curve of low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533). Rates of net and gross photosynthesis of low chlorophyll (white) and dark green parents (black) pods under field conditions. Each dot represents a value (n=4) {+/-}SE. We assumed that the seeds greatly inhibited the transmittance of light through the pod and used photosynthetic photon flux density for a single-side. C_FIG O_TBL View this table: org.highwire.dtl.DTLVardef@3f0528org.highwire.dtl.DTLVardef@16ba712org.highwire.dtl.DTLVardef@a5ab2aorg.highwire.dtl.DTLVardef@889254org.highwire.dtl.DTLVardef@3efa4f_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable 2.C_FLOATNO O_TABLECAPTIONPod photosynthetic parameters for low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533). Photosynthesis was measured 1 September through 15 September 2021 at the University of Illinois Energy Farm in Urbana, IL, USA. The statistical analysis was done using ANOVA with linear mixed model (alpha=0.05). N=4 {+/-} SEM for Clark and N=3 {+/-} SEM for Y11. C_TABLECAPTION C_TBL O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=130 SRC="FIGDIR/small/744892v1_fig2.gif" ALT="Figure 2"> View larger version (23K): org.highwire.dtl.DTLVardef@a36c26org.highwire.dtl.DTLVardef@1116c8forg.highwire.dtl.DTLVardef@ee5e61org.highwire.dtl.DTLVardef@1766712_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Low chlorophyll mutant (Y11/y11, PI547555) and its parent (Clark, PI548533) differ in concentration of seed oil, which interacts with height of pod and row spacing. The box plots show the median (central line), the lower and upper quartiles (box) and the minimum and maximum values (whiskers). The statistical analysis was done using ANOVA with linear mixed model (n=3 blocks, alpha=0.05). Least squares mean is used to compare. N.s., non- significant in the analysis. A. Concentration of oil in low chlorophyll mutant seeds from the upper canopy decreased by 4% compared to the dark green parent (18.2% vs 19%) while there was no difference between them in the seeds from the lower canopy (20.2% vs 20.6%). B. Schematic layout of 2013 field setting showing two different row spacings. C. Concentration of oil in low chlorophyll mutant decreased by 2% in 38cm spacing (21.4% vs 22%) while there was no difference in 19cm spacing (21.3% vs 21.7%) in 2013 field. C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=162 SRC="FIGDIR/small/744892v1_fig5.gif" ALT="Figure 5"> View larger version (22K): org.highwire.dtl.DTLVardef@68e508org.highwire.dtl.DTLVardef@94a6ccorg.highwire.dtl.DTLVardef@152a187org.highwire.dtl.DTLVardef@1eae137_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 5C_FLOATNO (greenhouse). Correlation between the level of leaf chlorophyll (x-axis: SPAD reading) and the level of immature pod or seed chlorophyll (y-axis, mg/g DW). Line represents the linear regression model. R-squared is a coefficient of determination, the percentage of the response variable variation that is explained by the linear model. Pod is labeled by the fresh weight of seeds it contained. A. Level of chlorophyll of 25-100mg pod (n=18). B. Level of chlorophyll of 100-200mg pod (n=17) . C. Level of chlorophyll of 25-100mg seed (n=17). D. Level of chlorophyll of 100-200mg seed (n=20). C_FIG O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=180 SRC="FIGDIR/small/744892v1_fig6.gif" ALT="Figure 6"> View larger version (28K): org.highwire.dtl.DTLVardef@167fd88org.highwire.dtl.DTLVardef@361472org.highwire.dtl.DTLVardef@786325org.highwire.dtl.DTLVardef@1b53855_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 6.C_FLOATNO Levels of gene expression in chlorophyll synthesis pathway. A. CHL common pathway genes; Glutamyl-tRNA reductase (GluTR). Glutamate 1- semialdehyde aminotransferase (GSA-AT). ALA dehydratase (ALAD). Uroporphyrinogen III synthase (UROS). Uroporphyrinogen III decarboxylase (UROD). Protoporphyrinogen IX oxidase (PPO). B. Mg branch; Mg-chelatase (Mgch). Magnesium-protoporphyrin IX monomethyl ester cyclase (MPEC). Protochlorophyllide reductase (POR). 3,8-divinyl protochlorophyllide a 8-vinyl-reductase (4VCR). Heme pathway; Ferrochelatase (FECH). Heme oxygenase (HO). Phytochromobilin synthase (HY). Data come from Severin et al (2010). RPKM, reads per kilobase per million mapped reads. DAF, days after flowering. The source seed is experimental line A81-356022 which was generated by introgressing G. soja (PI468916) into G. max (A81-356022). C_FIG

19
Comprehensive analysis of mulberry genetic diversity based on 1-DNJ content and SNP markers

Shen, Z.; Li, J.; Shi, J.; Li, Z.; Wang, F.; Geng, J.; Hu, K.

2026-08-19 genetics 10.64898/2026.08.11.744330 medRxiv
Top 0.6%
1.1%
Show abstract

Mulberry trees have high economic and ecological value, and a robust molecular marker system plus germplasm genetic diversity analysis is critical for innovative utilization of high-quality medicinal and economic mulberry germplasm. Here, 51 mulberry samples were used to develop SNP primers via genome resequencing, with the SNP-PCR system optimized by single-factor and orthogonal assays. The phenotypic diversity and SNP molecular marker genetic diversity of 1-deoxynojirimycin (1-DNJ) in mulberry leaves were analyzed respectively, and the genetic correlation between molecular markers and phenotypic traits was evaluated by Mantel test. Tested germplasm showed marked 1-DNJ variation (0.4805-2.5300 mg/g, CV=0.4241), reflecting rich genetic diversity. The optimal SNP-PCR system included Buffer (containing Mg{superscript 2}+) 2.2 L, 2.5 mM dNTP 0.4 L, forward and reverse primers (10 mol{middle dot}L-1) totaling 2.75 L, Taq DNA polymerase (5 U{middle dot}L-1) 0.3 L, DNA (50 ng{middle dot}L-1) 1.1 L, and ddH2O 13.65 L. 23 highly polymorphic ones amplified 91 loci (81 polymorphic, 89.10% polymorphism rate). Genetic diversity analysis showed that the average genetic distance was 0.3010, and the average expected heterozygosity (H) and Shannon information index (I) reached 0.4667 and 0.3104 respectively, indicating that the genetic differentiation among the tested mulberry germplasms was significant and the population had a moderate to upper level of genetic diversity. UPGMA clustering divided 51 germplasms into 6 major groups at a genetic similarity coefficient of about 0.7, while phenotypic clustering based on 1-DNJ content divided them into 2 major categories and 4 subcategories, with high 1-DNJ germplasm clustered independently. Mantel correlation analysis showed that 6 SNP sites were significantly weakly correlated with 1-DNJ content (r < 0.3, p < 0.05), and can be used as candidate molecular markers for subsequent genetic analysis of 1-DNJ content.This study established a stable mulberry SNP-PCR system, Analyze the molecular genetic characteristics of mulberry germplasm and DNJ phenotypic variation rules respectively, and provide basic data for cluster comparison. and provided a scientific basis for marker database improvement, germplasm identification and molecular-assisted breeding.

20
Root phenotypic plasticity improves yield stability when directed toward an adaptive integrated phenotype

Lopez-Valdivia, I.; Tawale, A. B.; Schierenbeck, M.; Sandoni, D.; Jones, D. H.; Kirschner, G. K.; Schneider, H. M.

2026-08-11 plant biology 10.64898/2026.08.10.744026 medRxiv
Top 0.6%
1.1%
Show abstract

Root phenotypic plasticity is often proposed to improve crop performance under stress, yet it remains unclear how much plasticity is beneficial and whether adaptive responses require changes across many traits or adjustments in few specific traits. Using public data of 6,500 field-grown maize and barley plants, this study examined the extent and distribution of root plasticity, and when it is associated with yield stability. We quantified root plasticity across nine anatomical and architectural traits using complementary statistical models and applied a feature-discovery framework to identify the drought-associated optimal integrated phenotypes and determine whether plasticity toward these phenotypes improved yield stability. More plasticity did not mean greater yield stability. Neither the number of plastic traits nor the magnitude of plastic responses predicted yield stability. Rather, we identified species-specific high-yielding, stable integrated phenotypes defined by distinct trait configurations. Critically, genotypes whose plastic responses moved their root phenotype toward these targets achieved greater yield stability, whereas movement away from them was associated with lower stability. Root plasticity is adaptive when it shifts root phenotypes towards an optimal integrated phenotype. These findings show that the value of plasticity depends on the trajectory of phenotypic change rather than its magnitude alone.